文章背景与核心概要
随着大语言模型和复杂AI推理工作负载的爆炸式增长,传统的以线程为中心的处理器架构在处理海量数据移动和并行计算时,逐渐面临能效比和可扩展性的瓶颈。为了突破这一限制,本文推出了Maia 200——一款采用软件定义局部访问数据流架构(SDLA)的先进AI加速器。
Maia 200的核心技术在于摒弃了传统的多线程控制模式,转而显式地对数据流引擎进行编程,以精准调度高度专用的内存和数据移动系统。该架构在750W的功耗(TDP)下,能够提供惊人的10,145 Tflop/s的FP4算力和5,072 Tflop/s的FP8算力,并具备7 TB/s的HBM带宽。这不仅极大地提升了AI推理的并行度和能效比,也为下一代高性能计算系统提供了一种兼具成本效益与扩展性的创新解决方案。
Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration
arXiv: 2608.24664 [cs.AR]
Submitted on: August 25, 2026
Primary Subject: Hardware Architecture (cs.AR)
Other Subjects: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
DOI: 10.48550/arXiv.2608.24664
arXiv: 2608.24664 [cs.AR]
Submitted on: August 25, 2026
Primary Subject: Hardware Architecture (cs.AR)
Other Subjects: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
DOI: 10.48550/arXiv.2608.24664
Authors
Authors
- Sherry Xu
- Marco Heddes
- Jackson Peng
- Tom Savell
- Monica Tang
- Prashant Ranjan
- Jesse Benson
- Ofer Dekel
- Saurabh Dighe
- Anupama Kurpad
- Artour Levin
- Matthew Mattina
- George Petre
- Cheng Tang
- Yuan Yu
- Li Zhang
- Torsten Hoefler
- Sherry Xu
- Marco Heddes
- Jackson Peng
- Tom Savell
- Monica Tang
- Prashant Ranjan
- Jesse Benson
- Ofer Dekel
- Saurabh Dighe
- Anupama Kurpad
- Artour Levin
- Matthew Mattina
- George Petre
- Cheng Tang
- Yuan Yu
- Li Zhang
- Torsten Hoefler
Summary
Summary
Maia 200通过从传统的以线程为中心的架构转向软件定义局部访问数据流架构(SDLA),开创了AI加速范式的转变。通过显式编程数据流引擎来协调专用的内存和数据移动系统,Maia 200为AI推理工作负载实现了巨大的并行性、高效率以及显著的成本和能耗节约。
Maia 200 introduces a paradigm shift in AI acceleration by moving away from traditional thread-centric architectures toward a Software Defined Locally Accessed Dataflow Architecture (SDLA). By explicitly programming dataflow engines to coordinate specialized memory and data movement systems, Maia 200 achieves massive parallelism, high efficiency, and significant cost and energy savings for AI inference workloads.
Key Performance Metrics
- FP4 性能: 10,145 Tflop/s
- FP8 性能: 5,072 Tflop/s
- 热设计功耗 (TDP): 750W
- HBM 带宽: 7 TB/s
Key Performance Metrics
- FP4 Performance: 10,145 Tflop/s
- FP8 Performance: 5,072 Tflop/s
- Thermal Design Power (TDP): 750W
- HBM Bandwidth: 7 TB/s
Abstract
我们推出了 Maia 200,这是一款先进的 AI 加速器,在 750W 的 TDP 和 7 TB/s 的 HBM 带宽下,能够提供 10,145 Tflop/s 的 FP4 算力和 5,072 Tflop/s 的 FP8 算力的高性能。Maia 堪称新一代软件定义局部访问数据流架构(SDLA)的典范,该架构显式编程数据流引擎以编排高度专用的内存和数据移动引擎。这种方法将焦点从当今以线程为中心的架构转移到以数据移动为中心的架构,从而提高了效率和可扩展性。我们受弗林分类法(Flynn's classification)启发的的数据管理分类法,凸显了 SDLA 如何应对现代 AI 计算中的挑战。Maia 200 在支持面向 AI 推理工作负载的大规模并行处理的同时,实现了显著的成本和能耗节约,使其成为下一代高性能计算系统的一个引人注目的解决方案。
We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate highly specialized memories and data movement engines. This approach shifts the focus from today's thread-centric to data-movement-centric architecture, improving efficiency and scalability. Our taxonomy of data management, inspired by Flynn's classification, highlights how SDLA addresses challenges in modern AI computing. Maia 200 achieves significant cost and energy savings while supporting massive parallelism for AI inference workloads, making it a compelling solution for next-generation high-performance computing systems.
Access Full-Text & Resources
- PDF 版本: 查看 PDF
- HTML 版本: HTML (实验性)
- TeX 源码: TeX 源码
- 许可证: 知识共享署名 4.0
Access Full-Text & Resources
- PDF Version: View PDF
- HTML Version: HTML (experimental)
- TeX Source: TeX Source
- License: Creative Commons Attribution 4.0
![]()
External References & Citations
External References & Citations